Papers with model pruning

14 papers
Iterative Structured Pruning for Large Language Models with Multi-Domain Calibration (2026.eacl-industry)

Copied to clipboard

Challenge: Existing models with unstructured pruning often yield irregular sparsity patterns that necessitate specialized hardware or software support.
Approach: They propose a structured pruning framework that eliminates entire architectural components and maintains compatibility with standard hardware accelerators.
Outcome: The proposed model pruning framework achieves significant compression with minimal performance degradation on multiple models across diverse downstream tasks.
Data Pruning for Efficient Model Pruning in Neural Machine Translation (2023.findings-emnlp)

Copied to clipboard

Challenge: Large-scale pre-trained language models have demonstrated encouraging performance in various NLP tasks at the cost of over-parametrized networks and high memory requirements.
Approach: They combine data pruning with movement pruning for Neural Machine Translation to enable efficient fine-pruning by leveraging cross-entropy scores of individual training instances.
Outcome: The proposed pruning strategy outperforms other pruning methods on a translation task and shows that training cross-entropy scores can reduce the steps required for convergence and training time.
BADGE: Speeding Up BERT Inference after Deployment via Block-wise Bypasses and Divergence-based Early Exiting (2023.acl-industry)

Copied to clipboard

Challenge: Recent years have witnessed the rise of many pre-trained language models (PLMs) such as GPT (Radford et al., 2019) and XLNet (Yang e.t al, 2019).
Approach: They propose a framework which consists of two off-the-shelf methods for improving PLMs’ early exiting.
Outcome: The proposed method can reduce the average latency of pre-trained language models and work with other inference speed-up methods like model pruning.
Efficient Contextualized Representation: Language Model Pruning for Sequence Labeling (D18-1)

Copied to clipboard

Challenge: Existing efforts to train pre-trained language models have brought significant improvements to various NLP applications.
Approach: They propose to compress bulky LMs while preserving useful information for a specific task.
Outcome: The proposed method can detach any layer without affecting others, and stretch shallow and wide LMs to be deep and narrow.
Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have found that arithmetic ability is limited to a few attention heads . existing studies do not elaborate on the mechanisms of these heads or how they influence FFN layers.
Approach: They propose a method that identifies an internal logic chain consisting of four stages from input to prediction.
Outcome: The proposed method improves prediction probabilities by amplifying coefficient scores of FFN neurons related to predictions.
Logits-Based Block Pruning with Affine Transformations for Large Language Models (2026.findings-eacl)

Copied to clipboard

Challenge: Existing methods for pruning models rely on calibration data and neglect cumulative effects of pruning on subsequent blocks.
Approach: They propose to use the Logit Disruption Score (LDS) to measure the impact of pruning by comparing the cosine similarity between the logits of the original and pruned models.
Outcome: Experiments on multiple datasets show that the proposed pruning technique reduces reliance on calibration data and improves generalization, achieving competitive results with existing methods.
LaCo: Large Language Model Pruning via Layer Collapse (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for model quantization, knowledge distillation, and model pruning are limited by hardware support limitations and the need for extensive training.
Approach: They propose a layer-wise structured pruner that collapses rear model layers into a prior layer and enables a rapid reduction in model size while preserving the model structure.
Outcome: The proposed pruner outperforms state-of-the-art pruning methods at pruning ratios of 25-30% and maintains an average task performance of over 80% at different pruning ratio.
Unveiling Multimodal Processing: Exploring Activation Patterns in Multimodal LLMs for Interpretability and Efficiency (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in multimodal large language models have remained opaque.
Approach: They propose a method to convert dense MLLMs into fine-grained Mixture-of-Experts architectures.
Outcome: The proposed method outperforms random expert pruning and sparse activation and model pruning.
BadWindtunnel: Defending Backdoor in High-noise Simulated Training with Confidence Variance (2025.findings-acl)

Copied to clipboard

Challenge: Current backdoor attack defenders in NLP typically involve data reduction or model pruning, risking losing crucial information.
Approach: They propose a backdoor defender that allows precise control over training conditions to model backdoor learning behavior without affecting the final model.
Outcome: The proposed model reduces the backdoor learning behavior without affecting the final model.
Let’s Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are composed of neurons that exhibit diverse behaviors and roles.
Approach: They propose a novel approach that refines the granularity of parameter training down to the individual neuron, enabling a more parameter-efficient fine-tuning model.
Outcome: The proposed approach exceeds the performance of full-parameter fine-tuning and PEFT and provides insights into the analysis of neurons.
Fisher Mask Nodes for Language Model Merging (2024.lrec-main)

Copied to clipboard

Challenge: Pre-trained models are ubiquitous in natural language processing, but individual fine-tuned models require significant overhead in multi-task scenarios.
Approach: They propose a method for fine-tuning pre-trained models for Transformers using Fisher information.
Outcome: The proposed method outperforms Fisher-weighted averaging in a fraction of the computational cost.
Time Course MechInterp: Analyzing the Evolution of Components and Knowledge in Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Large language models acquire and store factual knowledge for interpretability, reliability, efficiency . prior work on factual recall focused on localizing knowledge within transformer parameters .
Approach: They analyze the evolution of factual knowledge representation in a large language model by tracking its attention heads and feed forward networks over training.
Outcome: The proposed model acquires and stores factual knowledge over time and is adaptively trained . the proposed model can be pruned, optimized, and transparent .
Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their Applications (2024.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on how large language models process multiple languages, but internal mechanisms of LLMs remain insufficiently explored.
Approach: They propose to convert dense LLMs into fine-grained MoE architectures and analyze their activation patterns using expert activation frequency heatmaps.
Outcome: The proposed method outperforms random expert pruning and exceeds models in some languages.
PruMUX: Augmenting Data Multiplexing with Model Compression (2023.findings-acl)

Copied to clipboard

Challenge: Prior work has investigated methods like model pruning, knowledge distillation, and data multiplexing to increase model throughput without sacrificing accuracy.
Approach: They propose to combine structured pruning and data multiplexing methods to increase model throughput without sacrificing accuracy.
Outcome: The proposed method achieves 7.5-29.5X throughput improvement over a BERT-base model with accuracy threshold from 80% to 74%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations